Papers with medical image understanding
Can Medical Vision-Language Pre-training Succeed with Purely Synthetic Data? (2025.findings-acl)
Copied to clipboard
Che Liu, Zhongwei Wan, Haozhe Wang, Yinda Chen, Talha Qaiser, Chen Jin, Nikolay Burlutskiy, Fariba Yousefi, Rossella Arcucci
| Challenge: | Medical Vision-Language Pretraining (MedVLP) models typically require large-scale datasets with paired, high-quality image-text data. |
| Approach: | They propose to generate large-scale synthetic image-text pairs using off-the-shelf generative models . they propose to isolate model and training settings, focusing entirely from the data perspective. |
| Outcome: | The proposed pipeline outperforms models trained on real data by 3.8% on averaged AUC on zero-shot classification tasks. |
Self-Training Large Language and Vision Assistant for Medical Question Answering (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for collecting medical data are expensive and time-consuming. |
| Approach: | They propose a method to train a large-scale LVLM capable of auto-generating medical visual instruction data to improve data efficiency. |
| Outcome: | The proposed method shows that it performs well across three major visual question answering (VQA) benchmarks. |